Phase 2: Data & Mathematics Lesson 3 of 5

Vectors & Matrices:
Without the Pain

Linear algebra is the language AI speaks. Every image processed, every word understood, every recommendation made involves vectors and matrices. The good news: you do not need to memorise a single formula. You just need to see what they really are.

You will learn
What a vector actually is, explained in plain English
What a matrix is and why it matters
How images, text and data are all stored as vectors/matrices
What the dot product does and why it appears everywhere in AI

Why maths at all?

Before we go any further, let me address the question that every beginner has at this point: do I actually need to understand this?

The honest answer is yes, but not in the way you might fear. You do not need to perform matrix multiplication by hand on an exam. What you do need is a mental model for what is happening when your code runs. When you understand that an AI model is essentially doing an enormous number of vector operations, the behaviour of the model starts to make sense in a way it simply cannot if you treat it as a black box.

Here is the key insight: vectors and matrices are just ways of organising numbers. That is it. All the intimidating notation is just a shorthand for things you already understand.

"The miracle of the appropriateness of the language of mathematics for the formulation of the laws of physics is a wonderful gift which we neither understand nor deserve."

Eugene Wigner, physicist

A vector is just a list of numbers

Forget arrows. Forget geometry for now. At its most practical level, a vector is simply an ordered list of numbers. That is it. If you have ever looked at a row in a spreadsheet, you have worked with a vector.

A passenger (Titanic dataset)
3
22
7.25
0
1

This vector represents one Titanic passenger: ticket class 3, aged 22, fare £7.25, 0 siblings aboard, and 1 parent aboard.

Every row in your dataset becomes a vector. Every observation is a point in a multi-dimensional space, with one dimension per feature.

The number of elements in a vector is its dimension. The passenger vector above has 5 dimensions. A word embedding vector in a language model like GPT might have 768 or even 4,096 dimensions. You cannot visualise that space, but the maths works exactly the same way.

Think of it this way

Your location on Earth can be described by exactly two numbers: latitude and longitude. That is a 2-dimensional vector. If you add altitude, it becomes 3-dimensional. An AI model describing the meaning of a word might use 768 numbers, forming a 768-dimensional vector. More dimensions means more nuance, more information, more expressiveness.

A matrix is a table of numbers

A matrix is simply a grid of numbers arranged in rows and columns. When you look at your whole dataset, all your passengers and all their features together, you are looking at a matrix. Rows are observations. Columns are features.

5 passengers × 5 features: this dataset IS a matrix
Class
Age
Fare
Siblings
Parents
3
22
7.25
1
0
1
38
71.28
1
0
3
26
7.93
0
0
1
35
53.10
1
0
3
35
8.05
0
0

This is a 5×5 matrix. Shape: 5 rows by 5 columns. The highlighted row is one vector representing one passenger. When you call df.values in Pandas, you get exactly this.

Every matrix has a shape: (rows, columns). In NumPy you will write X.shape constantly.

Images are matrices

You saw in Lesson 2.1 that a grayscale image is a grid of numbers. That grid IS a matrix. A 28×28 pixel grayscale image is a 28×28 matrix, which is 784 numbers in total. A colour image adds a third dimension: 28×28×3 (one channel for red, green and blue). This 3-dimensional structure is called a tensor.

This is why the main AI framework is called TensorFlow and the operations are called tensor operations. Everything is just numbers in arrays of various shapes. The word "tensor" sounds intimidating but it is just a generalised word for a multi-dimensional array of numbers.

Shape matters enormously

In practical AI work, most bugs and errors come from mismatched shapes. You try to multiply two matrices but their dimensions do not align. You feed data with the wrong shape into a model. Understanding that every piece of data has a shape, and that shapes must be compatible for operations to work, will save you enormous frustration when you start coding.

Where vectors and matrices are used in AI

Vector
Word embeddings in language models
The word "king" is stored as a 768-dimensional vector. The word "queen" is a different vector, but mathematically close. This is how AI understands semantic similarity.
Vector
One row of your training data
Each observation in a dataset is a feature vector. The model learns to map these input vectors to an output label.
Matrix
The weights of a neural network layer
The learned parameters of every neural network layer are stored in weight matrices. Training updates these matrices. Inference multiplies your input by them.
Matrix
The attention scores in a transformer
Attention, the mechanism behind GPT and Claude, computes a matrix of scores showing how much each word should "attend" to every other word. This matrix is the heart of modern AI.

The dot product: AI's most important operation

There is one vector operation that appears more than any other in AI: the dot product. It takes two vectors of the same length and returns a single number. Here is why it matters.

Dot product: step by step
[2, 3, 1] · [4, 1, 5]
= (2×4) + (3×1) + (1×5)
= 8 + 3 + 5
= 16
Multiply each pair of matching elements, then add all the results. That single number captures how "aligned" or "similar" the two vectors are. High value = very similar direction. Near zero = unrelated.

This simple operation powers an enormous amount of AI. In neural networks, every layer computes a dot product between your data vector and a weight vector. In transformers, attention is computed using dot products between query and key vectors. In recommendation systems, similarity between a user vector and a product vector is computed with a dot product.

You do not need to compute this by hand. NumPy does it in one line: np.dot(a, b). But understanding what it means, asking "how similar are these two things?", is the insight that will help you understand why AI works the way it does.

Lesson Activity · No code required
Encode Yourself as a Vector
The best way to understand vectors is to create one from something real. In this exercise, you will turn yourself into a 5-dimensional vector. This is exactly what a machine learning model would do with you as a data point.
01 Choose 5 features about yourself that could be numbers: age, hours of sleep last night, number of apps on your phone, number of books read this year, daily steps yesterday.
02 Write them as a vector: [age, sleep_hours, app_count, books_read, daily_steps]. Then do the same for someone else in the group.
03 Compute the dot product of your two vectors by hand. The result is a single number. Does a higher number mean the two people are more "similar"? Why or why not? Discuss.
Your Notes
Studying independently? Write your thoughts or answers below. Notes save automatically to your browser.

Reflect

Before you move on

No right answers here. These questions are for you.

If word vectors place "king" close to "queen" and "man" close to "woman," what does it mean for two words to be "far apart" in vector space?

Distance in vector space represents semantic dissimilarity. Words used in similar contexts end up with similar vector coordinates, so words with unrelated meanings occupy distant regions of the space. "Far apart" does not mean spelling differences; it means the model rarely saw these words appear in the same kinds of sentences.

Matrix multiplication underlies almost every AI model. Why do GPUs handle this so much faster than CPUs, even though CPUs are generally more powerful per core?

CPUs have a small number of powerful cores optimised for sequential, complex logic. GPUs have thousands of simpler cores designed to run the same operation on many values simultaneously. Matrix multiplication is massively parallelisable: each output cell is an independent dot product. A GPU can compute thousands of those at once, while a CPU must largely do them one after another.

A colour image is stored as three matrices stacked together, one each for red, green, and blue. What does the number 0 mean in the red channel, and what does 255 mean?

0 in the red channel means that pixel contributes no red light at all. 255 is the maximum intensity, meaning full red. A pixel with red=255, green=0, blue=0 is pure red; red=255, green=255, blue=0 produces yellow. Every colour on a screen is a combination of these three channel values, and every pixel in an image is just three numbers.

Progress
Done with this lesson?
Mark it complete to track your progress.